Vision Research
○ Elsevier BV
Preprints posted in the last 30 days, ranked by how well they match Vision Research's content profile, based on 29 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Vlachou, M. E.; Thomas, E.; Blouin, J.
Show abstract
In this paper, we address the problem of quantifying similarity between planar 2D shapes, which is relevant to studies of internal representations in cognitive, developmental, and neurological research. We designed a set of test shapes arranged along a visually defined perceptual similarity gradient and used them to evaluate classical geometric methods for shape comparison, including Procrustes and Chamfer distance, as well as a convolutional neural network (CNN)-inspired feature-based method. Based on the limitations identified for these individual methods, we developed a hybrid Geometric-Feature Similarity (GFS) algorithm that combines geometric alignment, global contour properties, and convolutional feature-based descriptors into a unified weighted similarity score. By combining global geometric information with local structural features, the GFS algorithm more accurately reproduces human perceptual judgments of shape similarity than either geometric or feature-based methods alone. Requiring neither network training nor large labelled datasets, the proposed algorithm provides an efficient and interpretable tool for a broad range of studies involving quantitative shape comparison.
Pandey, P.; Pethe, S. R.; Indrajeet, I.; Ray, S.
Show abstract
Introduction: Decision making for selecting an object or a course of action from possible alternatives largely depends on our perceptual ability modulated by attention. When multiple stimuli appear close together in time, processing one stimulus can temporarily impair the processing of another due to temporal limitations of attention. Observers frequently fail to detect the second target (T2) presented within a few hundred milliseconds after the first target (T1) in a stream of stimuli, which is commonly known as attentional blink (AB). Existing theories attribute this perceptual lapse to T1 processing, distractor interference, or transient attentional gating; however, the computations underlying suppressive mechanism remains unresolved. We investigated whether pupil-size could reveal the underlying mechanisms of AB and predict conscious perception on a trial-by-trial basis. Methods: Pupil diameter and gaze locations were recorded using an infrared eye tracker. Machine learning techniques were used to classify trials when T2 was detected versus when it was not, after correct identification of T1, during an AB task from the pupil dynamics, which also yielded attentional episode (AE) associated with each element in the stream of visual stimuli when deconvolved. Results: Cross-validating classifiers achieved near-perfect accuracy not only in distinguishing but also predicting perceptual outcomes on a single-trial basis. AEs exhibited greater power when T2 was detected than when it was missed; the differential power in AEs on a logarithmic scale was highly synced with the differential pupil size. Conclusions: Collectively, these findings establish a framework for predicting attention-driven perceptual outcomes from pupil-dynamics at finer time-scale.
Spitschan, M.
Show abstract
PurposePupil diameter in daily life depends on both the light reaching the eye and the observers age, but established prediction formulas require laboratory quantities that are rarely measured in natural environments. We developed a compact age-corrected model that predicts pupil diameter from melanopic equivalent daylight illuminance (mEDI). MethodsWe used an existing field dataset in which binocular pupil diameter and near-corneal spectral irradiance were recorded while 83 adults aged 18-87 years moved through indoor and outdoor environments. The analysis included 10,082 valid paired observations. We fitted a bounded sigmoid relating pupil diameter to mEDI and age, with each participant given equal influence, and assessed prediction in participants excluded from model fitting. Performance was compared with simpler models, a flexible generalised additive model (GAM), and Watson-Yellott predictions based on assumed field geometry. ResultsPupil diameter decreased smoothly as mEDI increased. Age primarily reduced the difference between pupils in dim and bright conditions, by 0.768 mm per decade, while the predicted bright-light diameter changed little with age. In held-out participants, the bounded model had a participant-balanced root mean squared error (RMSE) of 0.630 mm and mean absolute error of 0.537 mm. The GAM had a slightly lower point-estimate RMSE of 0.610 mm, but the difference was small and uncertain. The bounded model outperformed the tested log-linear, reduced, age-only, and Watson-Yellott alternatives. ConclusionAge and mEDI are sufficient to provide useful population-average pupil predictions across the observed adult age and real-world light range. The model is transparent, physiologically bounded, and nearly as accurate as a flexible GAM, but predictions approaching darkness remain uncertain because valid mEDI measurements were not available in that range. Key pointsO_LIA compact equation predicts population-average pupil diameter from age and mEDI alone. C_LIO_LIAge mainly compresses the pupils response range by reducing pupil diameter under dimmer conditions. C_LIO_LIPrediction error in unseen participants was close to that of a flexible GAM, without requiring a fitted smooth object. C_LIO_LIThe model is intended for the observed adult age and field-light range, not for extrapolation into darkness. C_LI
de Jong, J.; Sergent, C.; Wexler, M.
Show abstract
The temporal resolution of vision is seriously limited. However, the response to a very brief flash, called the impulse response, is already quite sluggish at the earliest stages of vision, potentially obscuring the true temporal resolution of the rest of the visual system. Faster monitors that produce briefer flashes are subject to diminishing returns because, by definition, they cannot elicit responses that are any briefer than the impulse response. Here, taking inspiration from previous attempts, we develop a novel technique for presenting flashes that elicit 'briefer-than-brief' visual responses. Using a simple deconvolution technique, we reverse-engineer the visual response and estimate the form that the stimulus should take to elicit the response that a faster visual system would produce to a normal flash. Using psychophysics on human observers, we demonstrate that these 'briefer-than-brief' (BTB) flashes partially bypass the temporal limits presumably imposed by the early visual system using two paradigms: one that requires temporal segregation and one that requires temporal integration of sequential flashes. With BTB flashes, human observers successfully isolated two successive flashes at shorter intervals than with conventional flashes, improving temporal resolution by around 16%. We found that BTB stimuli not only improved temporal resolution, but also induced poorer performance on tasks requiring temporal integration, suggesting that the visual responses elicited by BTB flashes overlap less in time due to their briefer duration. In sum, our findings suggest that, using reverse-engineered stimuli, we can alleviate a temporal bottleneck that probably originates from the earliest stages of vision. In doing so, we allow higher visual areas to operate at a higher temporal resolution than previously thought possible.
Xiong, C.; Chen, Y.; Yang, Q.; Kim, S.; Meyyappan, S.; Bengson, J.; Mangun, R.; Ding, M.
Show abstract
Cueing paradigms are commonly used to study the neural mechanisms of visual spatial attention control. In these paradigms, each trial starts with an external cue, which instructs the subject to pay covert attention to a spatial location in anticipation of an impending stimulus (instructed attention). Recent work has introduced a new type of cue which prompts the subject to spontaneously decide which spatial location to attend (willed attention). We studied the neural mechanisms of willed attention control by analyzing fMRI and EEG data recorded at two institutions (UF and UC Davis) using the same willed attention paradigm. The findings include: (1) both instructional cues and the choice cue activated the DAN, (2) the choice cue additionally activated a frontoparietal decision network consisting of dorsal anterior cingulate cortex (dACC), anterior insula (AI), anterior prefrontal cortex (APFC), dorsal lateral prefrontal cortex (DLPFC), and inferior parietal lobule (IPL), (3) the decision about where to attend can be decoded in frontoparietal decision network in choice trials but not in instructional trials, and (4) EEG alpha oscillation patterns immediately preceding the choice cue, but not the instructional cues, predicted the postcue direction of attention and the frontoparietal decision network activity. Based on these findings we proposed a model of willed attention control suggesting how the direction of visual spatial attention was decided upon in the absence of external instructions.
Mardaljevic, J.; de Vries, S. W.; van Duijnhoven, J.
Show abstract
The measurement of light received at the cornea of the eye is a paramount consideration for the understanding of the relation between environmental illumination and the non-image-forming effects of light. The field of view (FOV) at the cornea is less than a full hemisphere, because it is partially occluded by human facial morphology. The International Commission on Illumination (CIE) has defined a standard model of human FOV. A suitably designed physical occluder attached to the sensor (of a light meter) has been proposed as a means of incorporating the effect of human FOV when taking measurements. Similarly, when using simulation to predict light received at the cornea, a geometrical description of the occluder at the eye point(s) can be added to the 3D model of the scene. The first occluder model proposed to represent CIE human FOV was enumerated in terms of: the CIE definition; the radius of the occluder; and, the radius of the light sensor disc. We present a simpler model based only on the CIE definition and the occluder radius. Both models were tested using a virtual goniophotometer. Various sensor response functions describing the spatial sensitivity across the sensor disc, including several we characterized through laboratory measurements, were included in the test. For all functions considered, the performance of the simpler occluder model was equivalent to or better than the model first proposed.
Zhang, J.; Liu, L.; Chen, J.; Zhao, N.; Li, H.; Yang, X.; Meng, X.; Ding, G.
Show abstract
Reading comprehension is a complex cognitive task that involves dynamic interactions between the brain and external information. Previous studies on reading development primarily focused on localized or static brain activities. However, it remains an enigma how brain state dynamics evolve with development underlying reading comprehension. This study aims to address this issue by combining functional magnetic resonance imaging (fMRI) with Hidden Markov Model (HMM) to explore brain state dynamics. A total of 35 typically developing children and 31 adults were scanned while reading a story. Our results demonstrated a tripartite brain state organization, characterized respectively by high activities in the visual (State #1), language (State #2), and default mode network (DMN, State #3) regions. Children exhibited significantly longer dwell time in the DMN state (State #3) compared to adults, along with a higher probability of transitioning from the language state (State #2) to the DMN state (State #3). In addition, adults exhibited greater flexibility in state transitions during reading comprehension. Finally, the alignment between the dynamic states of children and the average states of adults was a significant positive predictor of their reading comprehension performance. This study provides a novel, intuitive perspective on how brain state dynamics evolve during the development of reading comprehension.
Menetrey, M. Q.; Pascucci, D.
Show abstract
Several theories propose that perception and attention are governed by rhythmic processes that give rise to periodic fluctuations in behavior. However, empirical support for behavioral rhythms has been derived largely from paradigms involving brief, static stimuli. Here, we introduce a temporal averaging task requiring integration of rapidly unfolding visual features. Across three experiments, we tested averaging of orientation, size, and color under different eccentricity conditions. We used a temporally weighted averaging model to assess whether the influence of individual stimulus samples on perceptual estimates exhibits periodic modulation over time. We found no common rhythmic signature across tasks. Instead, orientation and size judgments showed reliable low-frequency modulations (<2.5 Hz), whereas color judgments showed only weak trends. Higher-frequency components (~3.5-8 Hz), often linked to theta and alpha rhythms, were observed only in a subset of participants and were limited to parafoveal orientation processing. These findings challenge the notion of universal behavioral rhythms and instead suggest that temporal dynamics are task-dependent, with slow oscillatory processes emerging as the most consistent feature.
Darjani, N.; Bakhtiari, S.; Vaziri-Pashkam, M.; Robert, S.
Show abstract
The human visual system integrates both static and dynamic information to support form and shape perception, yet the computational principles underlying the integration of motion for object recognition remain unclear. Artificial neural networks (ANNs) offer a computational framework for developing and testing hypotheses about these principles: if ANNs trained on motion-related tasks develop representations that align with brain activity and support object categorization, this would suggest that the training objectives and architectural constraints of these networks may capture key aspects of motion processing in biological visual systems in general, and motion processing for object recognition, in particular. Here, we investigated this question using "object kinematograms", stimuli in which object form is conveyed solely through motion cues. We measured neural responses of two higher regions of the lateral and the dorsal visual pathways, respectively, with strong sensitivity to dynamic cues from objects: lateral occipitotemporal cortex (LOTbio), and left supramarginal gyrus (SMGlh), as well as primary visual cortex (V1). We compared brain responses to representations extracted from two neural networks: SlowFast, a dual-pathway architecture trained on action recognition that processes slow- and fast-varying visual information with cross-pathway integration, and DorsalNet, a model of the primate dorsal visual pathway trained on embodied self-motion estimation. Representational similarity analysis revealed distinct representational profiles across brain areas, demonstrating functional specialization in motion-based form processing. LOTbio was best characterized by the slow pathway of the SlowFast model, whereas SMGlh showed strong similarity to both models. Critically, we found that representations aligned with brain activity also better supported behavioral function: the full SlowFast model, incorporating both slow and fast pathways, outperformed other models in few-shot categorization of object kinematograms and showed the highest similarity to human perceptual judgments. These findings demonstrate that with appropriate inductive biases, specifically, dual-pathway architectures for multi-scale motion processing and training objectives focused on dynamic visual tasks, ANNs can develop functionally useful representations of motion-defined forms that exhibit better alignment with the visual regions involved in processing dynamic visual signals.
Heirani Moghaddam, S.; Decarie, A.; Chua, R.; Cressman, E. K.
Show abstract
In the current experiment, we compared reported perceptual awareness of the visuomotor rotation to motor awareness of changes in reaches established using the process dissociation procedure and drawing task following visuomotor adaptation to a large (50 degrees; R50 group) or a small (30 degrees; R30 group) cursor rotation. Results revealed that perceptual and motor awareness did not differ in magnitude for the R50 group and were significantly correlated. In contrast, while the R30 group perceptually reported being aware of the visuomotor rotation, motor awareness was significantly less and responses were not significantly correlated across tasks. Overall, results suggest that perceptual and motor tasks assess different processes underlying visuomotor adaptation to a small cursor rotation, such that perceptual awareness of the visuomotor rotation is not reflected in reaching performance on tasks assessing motor awareness.
Chan, A. Y. C.; Shimojo, S.
Show abstract
This study characterizes how people combine visual and tactile directional cues while acting in a fully immersive 360{degrees} virtual environment. Participants used a vibrotactile belt and VR headset to localize targets while we manipulated visual reliability and the spatial discrepancy between visual and tactile signals. Behaviorally, degraded visual input made visual responses slower, less precise, and more susceptible to tactile pull, whereas tactile-guided responses remained comparatively stable. We then asked whether these behavioral changes reflected a change in multisensory binding or a change in sensory uncertainty. A Bayesian Causal Inference (BCI) framework captured the structure of behavior under high visual reliability and continued to track individual differences under low visual reliability, even though its absolute goodness-of-fit decreased. Under extreme visual noise, Bayesian Information Criterion sometimes favored a simpler Maximum Likelihood Estimation (MLE) model, but MLE showed poor absolute fit and did not capture meaningful behavioral variability. This dissociation shows that statistical parsimony and explanatory validity can diverge when behavior becomes highly variable. BCI-derived parameters further indicated that degraded vision increased visual uncertainty, while the prior tendency to bind visual and tactile cues remained stable. Kinematic analyses added a complementary insight: early movement trajectories were strongly shaped by tactile signals, even when final localization was visually guided. Together, these findings suggest that visual-tactile integration in 360{degrees} environments depends on sensory reliability and task demands, with tactile cues providing fast body-centered guidance when visual information is limited.
Yeh, L.-C.; Kaiser, D.
Show abstract
The attentional blink is a well-known phenomenon illustrating the limitations of human attention: When two visual targets are presented in rapid succession, identification of the second target is often impaired. While the attentional blink is known to attenuate when targets share perceptual features or category membership, real-world objects are also linked through contextual associations, shaped by objects typically occurring within the same environments. Here, we devised an attentional blink experiment in which we orthogonally manipulated contextual and categorical relationships between the two targets while controlling for their perceptual similarity. As the key result, contextual associations facilitated identification of the second target but impaired identification of the first target. These findings suggest that contextual associations yield distinct benefits and costs for visual cognition, where enhanced attentional access to subsequent targets is traded off against increased interference between targets.
Bai, X.; Kishimoto, K.; Sugiyama, O.; TAMURA, H.
Show abstract
This study aims to improve the detection performance of age-related macular degeneration (AMD) in low-quality retinal images. BackgroundAMD is a leading cause of vision loss among older adults globally, and accurate detection is crucial for clinical management. However, low-quality optical coherence tomography (OCT) images significantly compromise diagnostic accuracy. ObjectiveTo enhance AMD detection in low-quality images using noise-augmented data augmentation and an improved YOLO deep learning model. MethodsPublic datasets from UCSD and Duke University were utilized; the training dataset comprised 24,980 OCT images (high-quality and noise-augmented low-quality), while the testing dataset included 1,000 images (584 AMD, 416 normal). The model is based on the YOLOv8n framework, integrated with Squeeze-and-Excitation blocks (SEblock) and Adaptive Sparse Self-Attention (ASSA), with an additional 160x160 detection layer for detecting small lesions. Evaluation metrics included accuracy, sensitivity, specificity, and F2-score. ResultsThe proposed model achieved an accuracy of 99.02%, sensitivity of 98.17%, specificity of 100%, and an F2-score of 98.50% on the Duke dataset. Detection rates were significantly improved compared to traditional methods, particularly in low-quality images, with a detection rate of 89.60%, markedly superior to original YOLOv8n (55.10%) and classical models like ResNet50. ConclusionThe enhanced model, employing noise-augmented training data and improved attention mechanisms, demonstrates excellent AMD detection capabilities in low-quality OCT images, showing broad potential for clinical applications.
Dahech, H.; Minami, T.; Nakauchi, S.; Tamura, H.
Show abstract
Why does an angry face feel uncomfortable? The answer is that it signals a threat. However, a face is only part of an encounter, and distance, facial stimulus type, and gaze may shape discomfort regardless of perceived anger. To separate these cues, we conducted three within-subjects virtual reality experiments. In each experiment, 24 adults viewed avatars at intimate, personal, and social distances (30, 100, and 300 cm, respectively) and rated the faces perceived anger and their own discomfort; head movement was recorded in Experiments 2 and 3. In Experiment 1, the expression (angry, neutral) and facial color (natural, red) were crossed with distance; in Experiment 2, a featureless mannequin served as a nonface comparison; and in Experiment 3, the gaze direction (direct, averted) was manipulated. Expression primarily determined perceived anger, whereas distance predominantly determined discomfort: A nearby neutral face was uncomfortable despite low perceived anger (Experiment 1). A neutral human face was more uncomfortable than a mannequin, although both received similarly low perceived-anger ratings (Experiment 2). Direct gaze increased the discomfort without changing perceived anger (Experiment 3). Backward head movement exhibited a similar pattern, with participants leaning back more from human faces than from the mannequin. These results indicate that the discomfort associated with an angry face is not merely explained by perceived anger. Instead, social discomfort was differentially associated with interpersonal distance and gaze direction and differed between the human-face and mannequin conditions.
Sharifi Nowghabi, A.; Sharghilavan, S.; Bagheri, A.; Izadifar, M.
Show abstract
Wayfinding in hospitals is often hindered by ineffective signage; however, the cognitive mechanisms of healthcare wayfinding symbols comprehension remain under-researched. This study utilized eye-tracking and spatial gaze mapping to examine how visual complexity, abstraction, and human figuration modulate perception in 40 healthy adults viewing 24 hospital-related healthcare wayfinding symbols. Results indicate that pupil size is a sensitive physiological marker of cognitive load, significantly influenced by visual complexity ({chi}2 = 11.32, p = .022) and abstraction ({chi}2 = 7.49, p = .027). Human figuration reduced fixation duration and increased saccade amplitude, facilitating efficient semantic integration. Furthermore, human-centric healthcare wayfinding symbols elicited streamlined gaze trajectories, whereas abstract/complex designs induced chaotic scanpaths. These findings suggest that human figuration acts as a cognitive scaffold, reducing mental effort. We provide evidence-based guidelines for optimizing healthcare wayfinding symbols by prioritizing human body representations and balancing abstraction levels. HighlightO_LIPupil size indexes cognitive load during symbol comprehension. C_LIO_LIHuman figuration cuts fixation duration, boosting wayfinding efficiency. C_LIO_LIAbstract symbols increase pupil dilation, raising cognitive load. C_LI
Ha, L.; Sun, C.; Tang, R.
Show abstract
Analysis does not always enhance aesthetic experience. Philosophical accounts have long suggested that decomposing an aesthetic experience into determinate components may weaken it, yet this possibility has rarely been tested experimentally. To examine whether, when, and how analysis produces divergent effects on aesthetic experience, we conducted two experiments manipulating analysis depth. Experiment 1 showed that, during affective analysis of visual art, deep analysis produced a significantly weaker increase in aesthetic ratings than shallow analysis. In Experiment 2, we selected this condition to investigate the underlying mechanism. The behavioral effect was replicated: deep analysis removed the increase produced by shallow analysis without reducing ratings below the image baseline. Frequency-resolved brain network analysis further revealed a stronger task-related component and higher spatial entropy within the default mode network under deep analysis. Network-behavior correlations observed under shallow analysis were absent under deep analysis, suggesting reduced correspondence between the default-mode network (DMN) organization and aesthetic experience. Exploratory analyses further showed that spatial weights in the lateral temporal cortex and inferior parietal lobule were associated with smaller increases in aesthetic ratings. Together, these findings indicate that deeper analysis can selectively weaken improvements in aesthetic experience by altering how affective information is organized within the DMN.
Yildiran, O. F.; Ni, L.; Landy, M. S.
Show abstract
Previous work showed that observers integrate audiovisual duration cues optimally when cue-conflict is small. Does causal inference lead to a breakdown of audiovisual integration when duration conflicts are large? We addressed this by testing a wide range of duration cue-conflicts. Participants compared the auditory durations of a test and a standard stimulus. Audiovisual durations were consistent in the test stimulus, but differed by seven conflict durations (up to 250 ms) in the standard. Two levels of auditory noise were tested. Auditory duration percepts shifted systematically toward the visual duration, especially with high auditory noise. The shift was proportional to cue-conflict magnitude, inconsistent with causal inference. We compared several models. A heuristic model in which the observer probabilistically switches between the visual and auditory cues was preferred for most participants, although performance differences across models were small. Within the tested conflict range, the forced fusion, causal inference, and probabilistic cue switching models produced overlapping, near-linear shifts as a function of cue-conflict. Model simulations further revealed that given the measured sensory noise, forced fusion and causal inference can be discriminated only with unreasonably large conflicts. Together, while our results suggest that observers do not rely on causal inference when judging auditory durations under our conditions, high sensory encoding noise in auditory duration limits the discriminability of competing computational models.
Davies, T.; Bleeck, S.
Show abstract
Objective: This study investigated whether plosive consonants carry a perceptual loudness weighting that significantly exceeds that of non-plosive consonants when judged by hearing-impaired listeners. Design: A prospective loudness matching experiment utilizing the method of adjustment. Study Sample: 19 consenting native English speakers (Mean age: 61.4, SD: 16.4) with bilateral mild to moderate high-frequency sensorineural hearing loss, indicative of presbycusis. Stimuli: 13 vowel-consonant-vowel (VCV) nonsense syllables, exclusively utilizing the flanking vowel /u/. Results: Descriptive analysis revealed a strong time-order effect influencing loudness judgments for 7 of the 13 VCV test stimuli. Statistical testing showed no significant didference (P = 0.94) between the relative amplitudes corresponding to the point of equal loudness for plosive-containing versus non-plosive-containing VCV stimuli. However, 6 individual VCV stimuli, containing consonants from 4 separate manners of articulation, produced significant loudness matching data (P < 0.01). Conclusions: The results falsify the hypothesis that plosives, analyzed collectively as a class, possess a heavier perceptual loudness weighting than non-plosive consonants. While 6 individual VCV stimuli indicated potential individual consonantal loudness weightings, these findings must be interpreted cautiously due to the restriction to a single vowel context and the presence of procedural time-order biases.
Flieger, P.; Stecher, R.; Kaiser, D.
Show abstract
Humans rapidly assess the beauty of natural scene images. Previous EEG work suggests neural representations of beauty emerge early and are temporally sustained. Complementary fMRI work pinpoints the neural correlates of beauty to visual, frontal, and default-mode network areas. An integrated view of the spatiotemporal dynamics that give rise to the perception of beauty, however, is lacking. Beyond the beauty of the depicted scene, the quality of the image itself influences its perceived beauty, and it is unknown how the brain separates these two factors. To address these questions, we recorded EEG (N = 52) and fMRI (N = 29) data while participants rated the beauty of 100 natural scene photographs. Another group of participants (N = 46) rated the image quality of the same photographs. Separate representational similarity analyses on the EEG and fMRI data revealed early and sustained beauty-related representations across widespread cortical areas. In contrast, representations of image quality emerged earlier, had markedly different representational dynamics, and were predominantly localized to visual cortex. In a model-based EEG-fMRI fusion analysis, we investigated how the correspondence between temporally resolved EEG signals and spatially resolved fMRI signals is explained by beauty ratings. Our results suggest that beauty-related representations emerge early (from around 165ms and peaking at 275ms post-onset), are long-lasting, and primarily originate from high-level visual cortex. This spatiotemporal signature persisted when controlling for image-quality ratings. Our findings emphasize the importance of perceptual processing for perceived beauty and suggest that the brain represents aesthetic appeal independently of image quality.
Cortinovis, D.; Orlandi, G.; van Campenhout, L.; Bracci, S.
Show abstract
Recent work has revealed two food-selective areas in the lateral and ventromedial occipitotemporal cortex (OTC). These studies have shown that food selectivity in these regions cannot be explained by mid-level features like shape, texture, or colour but differences in their representational content remain unclear. Across two fMRI experiments conducted in the same group of participants, we characterized the dimensions underlying food representations in lateral and ventral OTC by examining the contribution of action-related object properties, such as manipulability, relevant to object interaction, and visual features, such as colour and ensemble statistics, relevant to object recognition. Our results reveal a clear dissociation between lateral and ventral OTC, indicating that food representations in these regions reflect distinct computational constraints. In lateral OTC, food representations were primarily associated with action-related properties shared between food and other graspable objects, whereas in ventral OTC, food representations were sensitive to surface object properties, such as colour and ensemble configuration. Consistent with this distinction, lateral OTC showed greater sensitivity to individual objects against distinctive background and responded equally to colour and greyscale stimuli, while ventral OTC exhibited greater sensitivity to coloured stimuli and ensembles with no distinctive background. Finally, topographic artificial neural networks implementing architectural constraints meant to capture OTC spatial organization similarly exhibited two dissociable clusters of food-selective units based on sensitivity to ensemble statistics. Together, these findings suggest that lateral food representations reflect action-relevant properties shared with other manipulable objects, whereas ventral food representations arise from surface-based visual features critical for food identification.